Dot products II
Lecture 17
Recap
$$ % Colors
% Coordinate vectors and matrices
% Common sets
% Abstract vector symbols
% Norms / absolute value
% Optional: dot product spacing (looks nicer in slides)
% Operators $$
Inner Product Space
- An inner product \(\left\langle \vec{u},\vec{v} \right\rangle\) (satisfying certain axioms) assigns a scalar to each pair of vectors \(\vec{u}\) and \(\vec{v}\).
- This scalar measures how “similar” or “correlated” the two vectors are.
- When we take the inner product of a vector with itself, \(\left\langle \vec{u},\vec{u} \right\rangle\) measures the “size” of the vector.
- The norm of \(\vec{u}\) is defined by \(\left\lVert \vec{u} \right\rVert = \sqrt{\left\langle \vec{u},\vec{u} \right\rangle}\).
Euclidean Vector Space with Dot Product
- The Euclidean space \(\mathbb{R}^n\) with the dot product is the most important example of an inner product space.
- In this setting, “correlation” and “norm” have geometric meanings: angles and lengths.
- \(\left\lVert \vec{u} \right\rVert^2 = \vec{u} \!\cdot\!\vec{u}\) (length from the dot product).
- \(\displaystyle \cos\theta = \frac{\vec{u} \!\cdot\!\vec{v}}{\left\lVert \vec{u} \right\rVert\,\left\lVert \vec{v} \right\rVert}\) (angle from the dot product).
Gram–Schmidt Process
- The Gram–Schmidt process is an algorithm that constructs an orthogonal (or orthonormal) basis for the vector space spanned by a linearly independent set \(S\).
- The idea: add vectors one at a time, making each new vector orthogonal to the previous ones.
- We do this by subtracting all components parallel to the earlier vectors, obtained using orthogonal projections.
Gram–Schmidt Process
- Let \(S=\{\vec{v}_1,\dots,\vec{v}_k\}\) be linearly independent and let \(W=\mathop{\mathrm{span}}(S)\).
- Set \(\vec{w}_1=\vec{v}_1\).
- For \(j\ge2\), remove all components parallel to the previous vectors: \[ \vec{w}_j = \vec{v}_j - \sum_{i=1}^{j-1} \operatorname{proj}_{\vec{w}_i}\vec{v}_j. \]
- Then \(\{\vec{w}_1,\dots,\vec{w}_k\}\) is an orthogonal basis of \(W\).
- Normalize each \(\vec{w}_j\) if an orthonormal basis is desired.
- Visualization :: Gram–Schmidt process in 3D
Real-World Example: MPG Analysis
- In statistics, covariance (a measure of how two random variables move together) is closely related to an inner product.
- Suppose we want to model a quantity that depends linearly on several variables.
- If the variables are correlated, it becomes difficult to isolate the effect of a single variable.
- Because correlated variables tend to move together, it is hard to determine which variable is responsible for changes in the response.
- The Gram–Schmidt process removes the shared movement by producing orthogonal (uncorrelated) directions.
- This allows us to measure the unique contribution of each variable.
- The following article illustrates this idea using automobile MPG data:
Decorrelating Features using the Gram–Schmidt Process
Matrix Transpose
Dot Product as Matrix Transpose
- Let \(\vec{u} = \langle u_1,\dots,u_n \rangle\) and \(\vec{v} = \langle v_1,\dots,v_n \rangle\).
- Recall that \(\vec{u} \!\cdot\!\vec{v} = u_1v_1 + \cdots + u_nv_n\).
- This equals the matrix product \[ [\,u_1\ \cdots\ u_n\,] \begin{bmatrix} v_1\\ \vdots\\ v_n \end{bmatrix}. \]
- The row matrix comes from writing a column vector as a row using the transpose.
Matrix Transpose
- We use the transpose to convert column vectors into row vectors: \[ \vec{u}^T = \begin{bmatrix} u_1\\ \vdots\\ u_n \end{bmatrix}^T = [\,u_1\ \cdots\ u_n\,]. \]
- In general, the transpose of an \(n\times k\) matrix \(A\) is a \(k\times n\) matrix \(A^T\).
- If \(A=[\vec{u}_1\ \cdots\ \vec{u}_k]\), then \(A^T=\begin{bmatrix}\vec{u}_1^T\\ \vdots\\ \vec{u}_k^T\end{bmatrix}\) in block notation.
- Entrywise definition: if \(A_{ij}\) is the entry in row \(i\), column \(j\), then
\((A^T)_{ij} = A_{ji}\) for \(1\le i \le k\) and \(1\le j\le n\).
Example
- Consider the \(3\times5\) matrix \[ A = \begin{bmatrix} 2 & -1 & 0 & 4 & 3 \\ 5 & 2 & -2 & 1 & 6 \\ -3 & 4 & 1 & 0 & -5 \end{bmatrix}. \]
- Find \(A^T\).
- What is \((A^T)_{21}\)?
Scan the QR code or go to join.iclicker.com/MBNJ.
Matrix Multiplication as Dot Products
- This viewpoint gives another way to understand matrix multiplication.
- When multiplying a matrix by a column vector, think of the matrix as a stack of row vectors.
- Let \(A^T=\begin{bmatrix}\vec{u}_1^T\\ \vdots\\ \vec{u}_m^T\end{bmatrix}\) where \(\vec{u}_i\in\mathbb{R}^k\), and let \(\vec{v}\in\mathbb{R}^k\).
- Then \[ A^T\vec{v} = \begin{bmatrix} \vec{u}_1\!\cdot\!\vec{v}\\ \vdots\\ \vec{u}_m\!\cdot\!\vec{v} \end{bmatrix}. \]
Example
- Let \[ A= \begin{bmatrix} 2 & -1 & 3\\ 0 & 4 & 1\\ -2 & 5 & 0\\ 1 & -3 & 2 \end{bmatrix}. \]
- Let \[ \vec{v}= \langle 1, -2, 3, 0 \rangle\in\mathbb{R}^4. \]
- Compute \(A^T\vec{v}\) using dot products.
General Matrix Multiplication
- Let \(\vec{u}_1,\dots,\vec{u}_n\) and \(\vec{v}_1,\dots,\vec{v}_m\) be vectors in \(\mathbb{R}^k\).
- Let \(A=[\vec{u}_1\ \cdots\ \vec{u}_n]\) be a \(k\times n\) matrix and \(B=[\vec{v}_1\ \cdots\ \vec{v}_m]\) a \(k\times m\) matrix.
- Then \(A^T\) is an \(n\times k\) matrix.
- The \((i,j)\) entry of \(A^T B\) is \[ (A^T B)_{ij} = \vec{u}_i \!\cdot\!\vec{v}_j. \] for \(1\le i\le n\) and \(1\le j\le m\).
- Thus matrix multiplication encodes all pairwise dot products between the columns of \(A\) and \(B\).
Example
- Let \[ A = \begin{bmatrix} 1 & 4 \\ -2 & 0 \\ 3 & -1 \end{bmatrix},\qquad B= \begin{bmatrix} 2 & 1\\ -1 & 3\\ 0 & 4 \end{bmatrix}. \]
- Compute \(A^T B\) using dot products.
Orthonormal Matrix
- Let \(S=\{\vec{u}_1,\dots,\vec{u}_n\}\) be an orthonormal basis of \(\mathbb{R}^n\).
- That is, \[ \vec{u}_i \!\cdot\!\vec{u}_j = \begin{cases} 1 & \text{if } i=j,\\ 0 & \text{if } i\ne j. \end{cases} \]
- Let \(P=[\vec{u}_1\ \cdots\ \vec{u}_n]\) be the \(n\times n\) matrix whose columns are these basis vectors.
- Then \(P^T P = I_n\).
- Conversely, if an \(n\times n\) matrix \(P\) satisfies \(P^T P = I_n\), then \(P\) is called an orthogonal matrix.
- Thus, an orthogonal matrix is precisely a matrix whose columns form an orthonormal basis of \(\mathbb{R}^n\).
Example
- The identity matrix \(I_n\) encodes the standard basis, which is orthonormal.
- Consider the matrix \[ P= \frac{1}{3} \begin{bmatrix} 2 & -2 & 1\\ 1 & 2 & -2\\ 2 & 1 & 2 \end{bmatrix}. \]
- Let its columns be \(\vec{u}_1,\vec{u}_2,\vec{u}_3\).
- Check that \(P^T P = I_3\), hence \(P\) is an orthogonal matrix.
- The set \(S=\{\vec{u}_1,\vec{u}_2,\vec{u}_3\}\) forms an orthonormal basis of \(\mathbb{R}^3\).
